Scan connections held as arrays of numbers - #21
Merged
Merged
Conversation
The scan read an object per connection and kept its state in objects keyed
by station code and trip id, so every step was a string-keyed lookup, every
query started from midnight, and every connection asked its trip's calendar
whether it ran. On the GB rail feed that is 115-175ms a query.
createTimetable numbers the stations as it is built and holds everything the
scan reads as parallel typed arrays: connections sorted by arrival with a
counting sort, which is linear and stable so connections arriving together
keep the order 2.0.0 scanned them in; footpaths grouped by origin station;
each trip's usable calls as one flat run. A scan's state is an array per
station and per trip, and journeys are built from the station and call
numbers, looking codes, trips and stop times up only for the legs returned.
A scan now:
- starts at the first connection arriving at or after the departure time,
found by binary search, since nothing arriving earlier could be boarded
- stops once a connection arrives after every destination has been reached.
Stopping on departure instead would end it at a connection that departs and
arrives in the same minute, before a trip arriving in that minute that the
passenger is aboard could take the tie
- asks each distinct calendar once per date rather than each trip per
connection
- counts a passenger aboard a trip from the earliest call one of its
connections has arrived at, not the call they boarded at, as 2.0.0 did by
station. A trip only picking up at a later call has carried nobody to it,
and treating it as boardable there let a tie build a change onto a train
that had already left
Benchmarked against 2.0.0 over the gb-transit feed published 11 September,
planning Tuesday 15 September, each in its own process and repeated:
2.0.0 this
build timetable 570-630ms 400-470ms
planner memory 195MB 129MB
32 standard queries, mean 115-117ms 2.8-3.5ms
268 coupling through trips, mean 130ms 1.4-1.7ms
400 random station pairs, mean 171-174ms 3.9-4.0ms
the same, a new date each query 177ms 5.2ms
All 1,100 answers are the same journeys 2.0.0 returns, trip for trip.
The API changes with it: loadTimetable and createTimetable replace loadGtfs
and toGtfsData, DepartAfterQuery takes the timetable and its filters,
ConnectionScanAlgorithm and JourneyFactory take the timetable and work in
station numbers, and ScanResultsFactory and the connection object types go.
Claude-Session: https://claude.ai/code/session_013JmX1aXaZ9tXrxkLxRhop4
The first version rewrote the scan as one function over a new Timetable, which was quick but no longer read like the algorithm or the code it replaced. This puts 2.0.0's structure back and changes only what the scan reads. toGtfsData, ConnectionScanAlgorithm asking ScanResults whether each connection isReachable and isBetter, setConnection, scanTransfers, isTransferBetter and setTransfer, and JourneyFactory's getLegs, toLeg, getCompactedLegs and getStopTimes are 2.0.0's, and DepartAfterQuery is unchanged. What they are given: - GtfsData numbers the stations in a StopTable, and holds Connections and Transfers as parallel typed arrays, interchange by station, and a TripCalendar that asks each distinct calendar once per date - a Connection is a number: a connection's index, or a footpath's below -1 - ScanResults holds arrays indexed by station and a trip arrival per trip, the earliest of its calls it has carried the passenger to - a scan starts at the first connection arriving after the departure time, and isFinished once a connection arrives after every destination was reached Every ScanResults method is inlined into the scan, so the structure costs little. Against the single function it replaces, standard queries and couplings take the same time and random pairs about 15% longer, while building is quicker and the timetable smaller, since no flat copy of every call is kept: 305-360ms and 83MB, against 400-470ms and 129MB. Against 2.0.0 the answers to all 1,100 benchmark queries are still the same journeys. Claude-Session: https://claude.ai/code/session_013JmX1aXaZ9tXrxkLxRhop4
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The algorithm and its classes are 2.0.0's. What they read is now arrays of numbers rather than an object per connection and objects keyed by station code, which makes queries 30–90× faster and the timetable less than half the size.
Unchanged
toGtfsData/loadGtfsConnectionScanAlgorithm.scanandscanTransfersScanResults:isReachable,isBetter,setConnection,isTransferBetter,setTransfer,isFinishedJourneyFactory:getLegs,toLeg,getCompactedLegs,getStopTimesDepartAfterQuery, untouchedWhat they read
StopTable.ConnectionsandTransfers: parallel typed arrays. Connections are sorted by arrival with a counting sort, which is stable, so ties keep 2.0.0's order.Connection: a number, the connection's index, or a footpath's index below -1.ScanResults: earliest arrivals and the connection index as arrays by station, plus a trip arrival per trip, the earliest call that trip has carried the passenger to.isFinishedonce a connection arrives after every destination was reached.TripCalendar: asks each distinct calendar once per date, not every connection.Every
ScanResultsmethod is inlined into the scan loop, so the structure costs little.Benchmark
GB rail feed published 11 September, planning Tuesday 15 September, each version in its own process, repeated:
All 1,100 answers are the same journeys 2.0.0 returns, trip for trip.
The first commit had a single-function rewrite that was about 15% quicker on random pairs. The second puts 2.0.0's structure back, and the numbers above are for that.
Keeping 2.0.0's answers
Diffing against 2.0.0 caught two mistakes in the first draft. Each now has specs that fail with the mistake put back:
isFinishedcompares a connection's arrival, not its departure. Otherwise a hop departing and arriving in the same minute ended the scan before the through train arriving that minute could win the tie.Breaking
new ConnectionScanAlgorithm(gtfs, new ScanResultsFactory(gtfs))andnew JourneyFactory(gtfs)GtfsDataholdsConnections,Transfers, interchange as anInt32Array, aTripCalendar, thetripsand astopTablescanreturns anInt32Arrayconnection index.TimetableConnectionandTransfersByOriginare gone.isTransferchecks a leg;isTransferConnectionchecks a connection.npm test(lint, typecheck, 54 specs),npm run buildandnpm pack --dry-runare clean.https://claude.ai/code/session_013JmX1aXaZ9tXrxkLxRhop4